在存在白噪声的情况下,在各个科学领域,在存在白噪声的情况下逃脱吸引盆地的平均退出时间至关重要。在这项工作中,我们提出了一种策略,以控制一般随机动力学系统的平均退出时间,以基于准潜电概念和机器学习实现所需的价值。具体而言,我们开发了一个神经网络体系结构来计算全局准次电位函数。然后,我们设计了一种系统的迭代数值算法来计算给定平均退出时间的控制器。此外,我们在有效的汉密尔顿 - 雅各比计划和受过训练的神经网络的帮助下确定了亚稳态吸引子之间的最可能路径。数值实验表明,我们的控制策略是有效且足够准确的。
为了获得下游图像信号过程(ISP)的高质量的原始图像,在本文中,我们提出了一个有效的本地乘法变压器,称为ELMFORMER,用于原始图像恢复。 Elmformer包含两个核心设计,尤其是针对原始属性是单渠道的原始图像。第一个设计是双向融合投影(BFP)模块,我们考虑了原始图像的颜色特征和单渠道的空间结构。第二个是我们提出了一个本地乘法自我注意力(L-MSA)方案,以有效地从当地空间传递信息到相关部分。 Elmformer可以有效地减少计算消耗,并在原始图像恢复任务上表现良好。通过这两种核心设计,Elmformer提高了最高的性能,并且与最先进的机构相比,原始DeNoising和原始Deblurring基准测试最低。广泛的实验证明了Elmformer的优势和概括能力。在SIDD基准测试中,我们的方法比基于ISP的方法具有更好的降解性能,这些方法需要大量的额外的SRGB培训图像。这些代码在https://github.com/leonmakise/elmformer上发布。
对人类流动性进行建模有助于了解人们如何访问资源并在城市中彼此进行身体接触,从而有助于各种应用,例如城市规划,流行病控制和基于位置的广告。下一个位置预测是单个人类移动性建模中的一项决定性任务,通常被视为序列建模,用Markov或基于RNN的方法解决。但是,现有模型几乎不关注单个旅行决策的逻辑和人口集体行为的可重复性。为此,我们提出了一个因果关系和空间约束的长期和短期学习者(CSLSL),以进行下一个位置预测。 CSLSL利用基于多任务学习的因果结构来明确对“ $ \ rightarrow $ wher wher wher wher whit $ \ rightarrow $ where where where”,a.k.a.”接下来,我们提出一个空间约束损失函数作为辅助任务,以确保旅行者目的地的预测和实际空间分布之间的一致性。此外,CSLSL采用了名为Long and Short-Charturer(LSC)的模块,以了解不同时间跨度的过渡规律。在三个现实世界数据集上进行的广泛实验表明,CSLSL的性能改善了基准,并确认引入因果关系和一致性约束的有效性。该实现可在https://github.com/urbanmobility/cslsl上获得。
Fairness and robustness are two important concerns for federated learning systems. In this work, we identify that robustness to data and model poisoning attacks and fairness, measured as the uniformity of performance across devices, are competing constraints in statistically heterogeneous networks. To address these constraints, we propose employing a simple, general framework for personalized federated learning, Ditto, that can inherently provide fairness and robustness benefits, and develop a scalable solver for it. Theoretically, we analyze the ability of Ditto to achieve fairness and robustness simultaneously on a class of linear problems. Empirically, across a suite of federated datasets, we show that Ditto not only achieves competitive performance relative to recent personalization methods, but also enables more accurate, robust, and fair models relative to state-of-the-art fair or robust baselines.
Unsupervised domain adaptation (UDA) for semantic segmentation is a promising task freeing people from heavy annotation work. However, domain discrepancies in low-level image statistics and high-level contexts compromise the segmentation performance over the target domain. A key idea to tackle this problem is to perform both image-level and feature-level adaptation jointly. Unfortunately, there is a lack of such unified approaches for UDA tasks in the existing literature. This paper proposes a novel UDA pipeline for semantic segmentation that unifies image-level and feature-level adaptation. Concretely, for image-level domain shifts, we propose a global photometric alignment module and a global texture alignment module that align images in the source and target domains in terms of image-level properties. For feature-level domain shifts, we perform global manifold alignment by projecting pixel features from both domains onto the feature manifold of the source domain; and we further regularize category centers in the source domain through a category-oriented triplet loss and perform target domain consistency regularization over augmented target domain images. Experimental results demonstrate that our pipeline significantly outperforms previous methods. In the commonly tested GTA5$\rightarrow$Cityscapes task, our proposed method using Deeplab V3+ as the backbone surpasses previous SOTA by 8%, achieving 58.2% in mIoU.
Different people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they still cannot generate diverse speaking styles in the final talking head videos. To tackle this problem, we propose a one-shot style-controllable talking face generation framework. In a nutshell, we aim to attain a speaking style from an arbitrary reference speaking video and then drive the one-shot portrait to speak with the reference speaking style and another piece of audio. Specifically, we first develop a style encoder to extract dynamic facial motion patterns of a style reference video and then encode them into a style code. Afterward, we introduce a style-controllable decoder to synthesize stylized facial animations from the speech content and style code. In order to integrate the reference speaking style into generated videos, we design a style-aware adaptive transformer, which enables the encoded style code to adjust the weights of the feed-forward layers accordingly. Thanks to the style-aware adaptation mechanism, the reference speaking style can be better embedded into synthesized videos during decoding. Extensive experiments demonstrate that our method is capable of generating talking head videos with diverse speaking styles from only one portrait image and an audio clip while achieving authentic visual effects. Project Page: https://github.com/FuxiVirtualHuman/styletalk.
